Journal of Medical Imaging
● SPIE-Intl Soc Optical Eng
Preprints posted in the last 30 days, ranked by how well they match Journal of Medical Imaging's content profile, based on 11 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Bennett, J.; Woodland, M.; Castelo, A.; Altaie, M.; Antony, A.; Siddiqi, N. S.; Long, J. P.; Brock, K. K.
Show abstract
Deep learning models deployed in clinical imaging frequently encounter distribution shifts, yet most out-of-distribution (OOD) detection methods are evaluated only on controlled research datasets. As a result, it is unclear whether existing approaches can reliably identify segmentation failures that arise in real-world clinical practice. We evaluated six OOD detection methods on a deployed liver CT segmentation model (3D nnU-Net) using internal data from 400 patients and external data from 100 patients collected across nearly 70 sites in 7 countries. One method was Pairwise Surface DSC, a surface-based extension of Pairwise DSC, that we introduced. OOD performance was measured using sensitivity, AUROC, and balanced accuracy, with thresholds determined on an independent cohort of 400 patients using the Youden J statistic. Statistical significance was assessed using McNemar tests and stratified bootstraps ( = 0.05) with Benjamini-Hochberg correction. Pairwise Surface DSC was the top-performing method, with perfect sensitivities (1.00), near-perfect AUROCs (0.97 internal; 1.00 external), and the highest balanced accuracies (0.94 internal; 0.88 external; p<0.001). These results show that automated failure detection for liver CT segmentation is clinically feasible and that Pairwise Surface DSC is a promising candidate for deployment. Our code is available at https://github.com/mckellwoodland/liver_ct_ood_translation.
Genske, U.; Laudani, A.; Yan, L.; Peng, Y.; Boening, G.; Ulas, S. T.; Wagner, M. P.; Diekhoff, T.; Hamm, B.; Jahnke, P.
Show abstract
Artificial intelligence (AI) applications in computed tomography (CT) imaging require objective and continuous testing, yet standardised methods for this purpose have not been established. Here, we present a framework using physical phantoms for standardised testing and monitoring of AI, demonstrated in liver lesion detection. We begin by designing phantoms tailored to the anatomical input domain expected by AI algorithms, and then systematically assess how AI performance is affected by variations in scanner technology and operation across two clinical CT systems. Next, we perform longitudinal monitoring, yielding consistent results over fifteen months on both systems. Finally, we validate clinical relevance by demonstrating that AI models trained on phantom data generalize effectively to patients and exhibit no evidence of phantom-specific adaptation. Our findings show that anatomically realistic phantoms enable standardised, site-specific testing and monitoring of AI, providing a proactive method for local and cross-institutional quality assurance.
Yamamoto, S.
Show abstract
CT perfusion (CTP) is central to acute-stroke and oncologic imaging, yet quantitative outputs vary substantially across vendor software, undermining reproducibility. We present an open, transparent core (ctp-core) that fits first-pass time-attenuation curves with a gamma-variate model, derives perfusion indices (peak enhancement, time-to-peak, bolus-arrival time, and area under the curve) analytically from the fitted parameters, and renders parametric maps with the ASIST-Japan standardized lookup table (a-LUT) so that visualization is comparable across sites. Every parameter, bound, and processing step is exposed. The method is validated on Monte-Carlo synthetic curves with known ground truth; no confidential or patient data are used. Across signal-to-noise ratio (SNR) levels 5 to 100 (200 independent runs per level) the pipeline recovers peak time to within 0.03-0.52 s and peak amplitude to within 0.4-8.1% (mean absolute error), degrading monotonically with noise; at a representative SNR of 20 it recovers peak time within 0.13 s, peak amplitude within 2.0%, and bolus-arrival time within 0.51 s, with fit quality R-squared = 0.98. The reproducibility demonstration is deterministic (fixed seed) and re-runs to bit-stable metrics. All code, the synthetic-data generator, the standardized-visualization module, evaluation scripts, and a 34-test suite are released openly for independent verification. The contribution is a fully open, parameter-transparent gamma-variate plus standardized-visualization pipeline with reproducible synthetic benchmarks: a reference others can audit, reuse, and build on.
Rich, J. M.; Kang, R.; Jin, D.; Subramanian, S.; Duddalwar, V.; Pachter, L.
Show abstract
We developed a standardized, reproducible preprocessing framework for computed tomography (CT) imaging data from multi-institutional repositories such The Cancer Imaging Archive (TCIA), enabling consistent radiomics and artificial intelligence (AI) analyses. Imaging data from TCGA-KIRC patients available on TCIA were used as a representative heterogeneous dataset characterized by variation in acquisition protocols, inconsistent metadata, and differing image quality. The proposed modular pipeline includes series filtering, DICOM-to-NIfTI conversion, orientation harmonization to a canonical coordinate system, voxel spacing normalization, intensity clipping and normalization, segmentation integration, and metadata validation, and is implemented in a reproducible, notebook-based framework compatible with common radiomics and deep learning workflows. This pipeline standardizes imaging data into analysis-ready volumes with consistent geometry, intensity distributions, and spatial alignment, reducing non-biological variability that can adversely affect radiomic feature stability and model performance. The modular design enables task-specific adaptation of individual preprocessing steps while maintaining overall consistency. Although demonstrated on TCIA, this framework is generalizable to other heterogeneous imaging datasets and provides a foundation for robust, large-scale computational imaging studies.
Pasyar, P.; Mei, K.; Im, J. Y.; Roshkovan, L.; Geagan, M.; Noël, P. B.
Show abstract
ABSTRACT Background: Metallic implants such as orthopedic screws, prostheses, and dental hardware produce beam-hardening, photon-starvation, and streak artifacts that degrade computed tomography (CT) image quality, and the metal artifact reduction (MAR) methods developed to mitigate them require objective, reproducible benchmarking. Purpose: Objective evaluation of MAR algorithms in CT is hindered by the absence of phantoms that simultaneously provide anatomically realistic backgrounds, embedded implants of known geometry, and controllable, ground-truth--referenced artifact intensity. We present a dual-filament, voxel-level three-dimensional (3D) printing method that fulfills these requirements and demonstrate its capabilities on a clinically representative cervical spine case with embedded orthopedic spinal screws. Methods: The proposed method extends the PixelPrint framework, a fused-deposition-modeling (FDM) workflow that converts clinical Digital Imaging and Communications in Medicine (DICOM) data directly into 3D-printer Geometric code (G-code) without intermediate segmentation or surface meshing, to interleaved, voxel-level deposition of two filaments: a calcium-doped polylactic acid (PLA) for soft tissue and bone, and a higher-attenuation metal-doped PLA for metallic implants. For demonstration, anonymized DICOM data of a healthy cervical spine were used to design and fabricate three matched phantoms, each with six embedded spinal screws at C4--C6: a 0% metal-infill ground-truth phantom, a 50% medium-metal-infill phantom, and an 85% high-metal-infill phantom. All phantoms were scanned on a clinical spectral CT system at 120 kVp and 1000 mAs, reconstructed at 0.67 mm slice thickness with virtual monoenergetic imaging (VMI) across 50--190 keV. Method performance was characterized by region of interest (ROI)-based Hounsfield Unit (HU) agreement with the source patient data and by the noise-independent Gumbel-distribution p-index metric. Results: The dual-filament method reproduced patient anatomy, soft-tissue contrast, and screw geometry with high fidelity. ROI HU values agreed with patient data within {+/-}25 HU for soft tissue and trabecular bone; cortical regions were underestimated owing to the current ceiling of the calcium-doped PLA used in this study. The tunable-artifact behavior was quantified as follows: the Gumbel location parameter scaled monotonically from 46.7 HU (no-metal background) to 57.1 HU (50% infill) to 90.5 HU (85% infill) for the VMI 70 keV with standard filter. High-keV VMI reconstructions substantially reduced streak and beam-hardening artifacts while preserving anatomic detail. Conclusions: The proposed dual-filament, voxel-level PixelPrint method enables the fabrication of patient-specific, multi-material CT phantoms with embedded metallic implants and controllable, ground-truth--referenced artifact intensity. Although demonstrated here in a single cervical-spine case, the workflow is anatomy- and implant-agnostic by construction and could in principle be adapted to other musculoskeletal sites (e.g., knee, hip, dental) and implant materials, providing a reproducible methodological foundation for benchmarking MAR algorithms, characterizing spectral CT performance, and validating emerging photon-counting detector systems. Keywords: 3D printing methodology; fused deposition modeling; voxel-level multi-material printing; spectral computed tomography; metal artifact reduction; phantom design; orthopedic implants; dual filament; PixelPrint.
Chen, W.-Y.; Wan, S.-Y.; Lin, G.-Y.
Show abstract
Accurate segmentation of thin-wall organs-at-risk (OARs)-the cochlea, vestibular semicircular canals, internal auditory canal, tympanic cavity, and middle ear-is clinically relevant for head-and-neck radiotherapy planning, yet these small, thin-wall structures remain among the most challenging targets for automated delineation. Dual-frequency feature fusion is a promising direction for boundary-sensitive representation, but under the investigated FP16 FFT-FcaNet setting, we observe an approximately 863-fold activation-scale mismatch between the FFT and FcaNet branches, causing a nominal 5 percent residual coefficient to behave as an approximately 43-fold dominant term. We propose FreqFuseNet, which resolves this mismatch by normalizing the FcaNet branch to the FFT activation scale before residual injection with a fixed low-amplitude coefficient (beta = 0.05), restoring beta as an interpretable 5 percent residual-amplitude coefficient relative to the FFT feature scale. Under a controlled binary per-OAR ROI protocol on the SegRap2023 head-and-neck CT benchmark across 10 clinically prioritized thin-wall OARs, FreqFuseNet achieves Dice of 0.849, HD95 of 0.824 mm, and SDice@1mm of 0.959 in the primary seed, with comparable performance in an independent second seed (Dice 0.843, HD95 0.823 mm). FreqFuseNet yields statistically significant case-level aggregate improvements over 3D U-Net and MedNeXt-S (Wilcoxon p < 0.01 and p < 0.05, respectively), using only 29.7 million parameters versus 414.6 million for the full wavelet baseline.
Gu, X.; Zhu, H.; Zhong, F.; Teng, G.-J.
Show abstract
Background: Nuclear medicine and radiopharmaceutical development require coordinated radiochemistry, dosimetry, molecular imaging, radiation-safety and clinical decision processes. Current workflows remain fragmented, difficult to audit and poorly standardised for evaluating domain-specific AI support. Methods: We developed RadGuide AI, a nuclear medicine agent built around a traceable data-model-tool loop. Patent, literature and clinical-trial records were converted into 15,596 initial QA items; relevance screening, completeness checks, semantic deduplication and cross-validation retained 5,474 core QA items. MedGemma-27B-Instruct served as the foundation model and was adapted with LoRA. The system incorporated 55 MCP-wrapped tools covering radiopharmaceutical R&D, clinical decision support, imaging analysis and radiation-safety/dosimetry. Evaluation used a locked N=200 benchmark with predefined denominators, leakage control, expert scoring, statistical procedures, factuality audits and tool-execution metrics. Results: RadGuide-LLM achieved 88.5% answer accuracy (177/200; 95% CI, 83.3-92.2%) and a Macro-Average score of 21.5/25 (bootstrap 95% CI, 20.9-22.0), exceeding GPT-4o, DeepSeek-V3.2 and the base MedGemma model in this technical evaluation. Supplementary audits reported guideline compliance, terminology recall, knowledge coverage, tool-routing success and preclinical/phantom dosimetry agreement with explicit denominators and confidence intervals. Interpretation: RadGuide AI converts nuclear medicine queries into auditable retrieval, tool selection, calculation, verification and reporting workflows. The findings support technical feasibility, not definitive patient-level clinical validation; prospective multicentre studies and external benchmark release remain required before clinical deployment.
Noyan, H.; Hickstein, R.; Ammann, C.; Kuhnt, J.; Fenski, M.; Prieto, C.; Botnar, R. M.; Hadler, T.; Hickstein, C.; Daud, E.; Blaszczyk, E.; Groeschel, J.; Lim, C.; Schulz-Menger, J.
Show abstract
Background: Epicardial adipose tissue (EAT) is a metabolically active fat depot adjacent to the myocardium and the coronary arteries that can be non-invasively assessed by cardiovascular magnetic resonance (CMR). Increased EAT volume quantified by CMR has been linked to adverse cardiac remodeling, atrial fibrillation, coronary artery disease, and heart failure. Among CMR techniques, isotropic three-dimensional (3D) Dixon imaging at 1.3 x 1.3 x 1.3 mm3 resolution was developed to improve tissue characterization, providing fat-water signal separation for precise volumetric EAT assessment. However, manual segmentation of 3D datasets is highly time-consuming. For integration into clinical and research CMR workflows, reliable and fast automated segmentation is needed. Purpose: To develop and evaluate an automated deep-learning-based pipeline for ventricular EAT quantification based on isotropic 3D Dixon CMR acquisitions. Methods: An nnU-Net model was trained on 165 3D Dixon CMR cases encompassing healthy individuals and patients with underlying cardiovascular disease. The model was trained using all four Dixon phase images (opposed-phase, in-phase, fat-phase, water-phase). Manual 3D ventricular EAT segmentations served as the ground truth for training and evaluation. Performance was evaluated in 30 independent cases using Dice similarity coefficient (DSC), 95th percentile Hausdorff distance (HD95), volumetric agreement, Pearson correlation, intraclass correlation (ICC), and Bland-Altman analysis. Model performance was benchmarked against interobserver and intraobserver variability. Results: Automated segmentation achieved a mean DSC of 0.896 {+/-} 0.039 and HD95 of 1.84 {+/-} 0.93 mm versus ground truth. Volumetric agreement with ground truth was high (r = 0.984, ICC = 0.988, p < 0.001; mean bias -0.70 mL, limits of agreement (LoA) [-10.31, 8.90] mL), exceeding interobserver agreement (bias -25.24 mL, LoA [-42.81, -7.66] mL) and comparable to intraobserver reproducibility (bias 2.72 mL, LoA [-8.73, 14.17] mL). Automated segmentation required less than one minute per case compared to 58.4 {+/-} 7.9 minutes for manual segmentation. Two of 30 cases (6.7%) required minor manual correction, both less than five minutes. Conclusion: Fully automated nnU-Net-based ventricular EAT segmentation from isotropic 3D Dixon CMR achieves accuracy comparable to intraobserver reproducibility while significantly reducing post-processing time. The approach may facilitate large-scale and longitudinal EAT quantification in CMR-based research workflows.
Amiri, S.; Afshar, P.; Rohban, M. H.
Show abstract
Objectives. Radiomics pipelines extract hundreds of quantitative features that are widely known to be redundant, but the structure of this redundancy is usually treated as a per-dataset nuisance to be pruned away. We tested the alternative hypothesis that a substantial number of feature-feature correlations are universal: they persist across patients and across anatomically distinct structures because they reflect shared mathematical and image-statistical properties of how the image is summarised, rather than properties of the tissue being imaged. Materials and Methods. We re-analysed the publicly available Radiomics Atlas Dataset of normal Abdominal and Pelvic CT (RADAPT), restricting the analysis to the 526 non-contrast-enhanced examinations of the 531-subject atlas and to the 107 original (non-filtered) PyRadiomics features. The 53 segmented structures were grouped into four broad anatomical categories -- bones, muscles, vessels, and parenchymal organs. RADAPT is distributed as one Excel file per structure, with patients as rows and features as columns. Within each structure file we z-score-normalised every feature across patients, computed the absolute Spearman correlation matrix, and retained edges with |{rho}| [≥] {tau} for {tau} in {0.70, 0.80, 0.90}. We then intersected the edge sets across all structure files to obtain a "universal" correlation graph, in which an edge survives only if it exceeds the threshold in every structure (each estimated across the full patient sample). Stable feature communities were defined as the maximal cliques of this graph. Robustness to patient sampling was tested by repeating the entire pipeline on five independent random splits of each file into two patient halves (10 sub-cohorts per threshold), and the implementation was independently reproduced in R. Results. Despite the strictness of the global-intersection criterion, 34, 24, and 14 stable feature communities survived at {tau} = 0.70, 0.80, and 0.90 respectively, with the largest cliques containing six members at {tau} = 0.70 and {tau} = 0.80 and five members at {tau} = 0.90. The community structure was clearly interpretable: separate cliques captured (i) variance-like intensity dispersion, (ii) long-run / low-frequency (coarse) texture, (iii) high gray-level texture, (iv) low gray-level texture, (v) volume and surface shape, and (vi) local-homogeneity and energy/entropy duals. On random-half resampling the exact-match recovery rate of these communities was 81.5 %, 86.7 %, and 80.7 % across the three thresholds; departures from exact recovery were almost always a single boundary feature added or dropped, consistent with finite-sample fluctuation of near-threshold edges rather than structural instability. The R re-implementation reproduced the Python results exactly. Conclusion. A substantial portion of radiomics feature collinearity is universal across patients and tissues. We distinguish two layers within it: trivial near-algebraic duals that are universal by construction, and non-trivial cross-matrix-family communities that are the genuine empirical finding. Together they provide an interpretable, definition-grounded basis for aggressive dimensionality reduction, for retrospectively reconciling apparently different feature selections in the literature, and for moving radiomics pipelines toward organ-agnostic, more reproducible models. Clinical relevance statement. Selecting a single representative feature from each universal community shrinks the original-feature space by roughly an order of magnitude without sacrificing biologically distinct information. For example, the five variance-family members (first-order Variance, GLCM SumSquares, GLCM ClusterTendency, GLDM and GLRLM GrayLevelVariance) can be replaced by a single representative, removing redundant degrees of freedom that would otherwise inflate model variance; and labelling each retained feature by its community lets two studies that selected different variance-family names be recognised as having found the same signal, simplifying model development and improving cross-cohort generalisability in clinical CT workflows.
Shanbhag, A.; Miller, R. J.; Killekar, A.; Marcinkiewicz, A. M.; Zhou, J.; Lemley, M.; Kamagate, A.; Van Kriekinge, S. D.; Kavanagh, P. B.; Feher, A.; Miller, E. J.; Liang, J. X.; Berman, D. S.; Dey, D.; Leahy, R. M.; Slomka, P.
Show abstract
Background: Coronary artery calcium (CAC) is an established measure of coronary atherosclerosis from computed tomography (CT). While deep learning (DL) can quantify CAC from non-dedicated CT, the accuracy is limited by image quality. Purpose: We derived and validated a novel method for DL CAC segmentation on ultra-low dose CT attenuation correction (CTAC) scans that is trained with synthetic low-dose, ungated images. Materials and Methods: Models were trained using one center and externally tested in two other centers. Synthetic, ungated CT scans were generated so that expert segmentations from dedicated CAC scans could be used as ground truth for perfectly registered synthetic images through knowledge adaptation (KAD-CAC). We evaluated agreement between CAC scoring methods vs expert readers on a per-patient and per-vessel basis, as well as associations with the primary outcome of death or myocardial infarction (MI). Results: The DL models were externally tested on 5969 patients with a median age of 64 (IQR 56 - 73), of whom 50.2% were male. The KAD-CAC model had higher Cohens kappa K (0.86, 95% CI 0.85 - 0.87) compared to previous convolutional LSTM model (K 0.78, 95% CI 0.76 - 0.80, p<0.01), or models trained with only gated images (K 0.81, 95% CI 0.80 - 0.82, p<0.01). Net reclassification improvement for CAC stratified risk of death or MI, was greatest for the KAD-CAC model over baseline including age, sex, hypertension, diabetes, dyslipidemia, family history, smoking, stress total perfusion deficit, and left ventricular ejection fraction. Conclusion: We use paired synthetic ungated scans to transfer expert gated CAC annotations into the ungated domain, resulting in substantially better vessel-level CAC scoring and improved risk stratification.
Hamkins, H. M.; Tam, K. H.; Sobremonte, A.; Jogi, S.; Koay, E.; Hassanzadeh, C.; Segars, P.; Tyagi, N.; Subashi, E.
Show abstract
Background: Independent end-to-end verification of adaptive radiotherapy on MR-Linac systems is limited by the lack of patient-specific phantoms able to reproduce imaging and dosimetric properties from CT and MRI scanners. We present a method for automated generation of 4D, patient-specific, multi-material 3D-printable phantoms for quality assurance of adaptive radiotherapy on a 1.5T MR-Linac. Methods: Patient images were automatically segmented using a pretrained deep learning model. The segmented structures were converted into high-resolution 3D meshes and assembled into printable phantoms. A dosimeter holder was inserted at user-defined anatomical locations, with orientation optimized to avoid traversal across heterogeneous tissue interfaces. Physiological motion was incorporated by generating phantoms from images at different timepoints and interpolating deformation fields to create continuous 4D models. Multi-material organs designed by mixing a set of six polymers at various proportions were used to reproduce tissue-specific imaging properties. The properties of material mixtures were evaluated in a clinical CT simulator and a 1.5T MR-Linac. Results: The proposed workflow enables automated generation of anatomically realistic phantoms with several types of embedded dosimeters. A discrete search method was designed for placement and immobilization of OSLD, film, and ion chamber dosimeters. Calibration curves for Hounsfield units were derived through variations in radiopaque material content, while MR signal intensity was modulated by gel and tissue matrix mixtures. Patient-derived abdominal phantoms were fabricated at multiple scales while replicating internal anatomical detail. Multi-dimensional phantom generation enabled continuous representation of motion states with consistent mesh topology across phases. Conclusions: We demonstrate an end-to-end workflow for automated generation of 4D patient-specific phantoms for MR-Linac quality assurance. The method combines realistic anatomy, embedded dosimetry, multimodal imaging properties, and physiological motion within a single fabrication framework. This approachmay enable an improved validation of adaptive radiotherapy workflows in MR-guided treatment devices.
Bazhutina, A.; Chumarnaya, T.; Zubarev, S.; Budanova, M.; Stepanova, V.; Khamzin, S.; Lebedev, D.; Solovyova, O.
Show abstract
Background: Cardiac resynchronization therapy (CRT) fails in 30% of patients, often due to suboptimal left ventricular pacing site (LVPS) selection. Current practice lacks tools for pre-procedural, patient-specific LVPS optimization within the accessible coronary sinus (CS) tributaries. This study aimed to develop a digital twin and an explainable ML-based clinical decision support framework to address this issue. Methods: Personalized 3D cardiac models incorporating ventricular anatomy, myocardial fibrosis, and CS anatomy were constructed from CT and LGE-MRI for 74 CRT candidates. Finite-element Eikonal simulations of biventricular pacing generated patient-specific electrophysiological features at candidate LVPS. A Machine Learning (ML) classifier was trained on a hybrid feature set of pre-procedural clinical variables and model-derived indices, validated by leave-one-out cross-validation. SHAP analysis provided a physiologically interpretable rationale for each prediction. The framework was applied to a pilot cohort of 19 patients with reconstructed 3D CS anatomy to generate a spatial likelihood map of CRT response across all clinically implantable pacing sites within each patient's CS. Results: The ML classifier outperformed the reference Feeny clinical calculator under LOO-CV (accuracy 0.78 vs 0.58; F1-score 0.75 vs 0.43), AUC=0.78, sensitivity=0.80, specificity=0.77. Bootstrap analysis yielded mean AUC=0.85 (95% CI 0.70-0.95). In the pilot CS cohort, the framework identified that 8 of 13 clinical non-responders had no accessible CS site predicted to yield a positive response, supporting redirection towards alternative pacing strategies. In the remaining 5, alternative implantable sites with high predicted response probability were identified. SHAP analysis confirmed that dominant predictors were patient-specific in their relative contributions, supporting individualized over heuristic-based LVPS selection. Conclusion: This pilot study demonstrates the feasibility of a digital twin and explainable ML framework as a pre-procedural clinical decision support tool for CRT planning, stratifying patients and identifying optimal implantable sites with transparent anatomical rationale. Prospective validation and regulatory evaluation are required before clinical deployment.
Mahtabi, B.; Nasr-Esfahani, E.; Yaraghi, S.
Show abstract
Pneumonia is a leading cause of infectious disease mortality worldwide, accounting for approximately 2.5 million deaths annually and 15% of deaths in children under five. Chest X-ray imaging remains the primary diagnostic tool, but accurate interpretation requires radiological expertise that is disproportionately concentrated in high-income settings, creating a diagnostic gap where disease burden is highest. Automated deep learning offers a scalable complement to specialist-dependent diagnosis, yet clinical adoption requires both high accuracy and transparent, interpretable reasoning. Convolutional neural networks (CNNs) have shown strong potential for pneumonia detection from chest X-rays, but two barriers impede clinical translation: the interpretability of black-box models and the computational feasibility of large architectures in resource-constrained settings. Explainable AI (XAI) methods such as Grad-CAM, Grad-CAM++, and Score-CAM address the interpretability barrier, yet systematic quantitative comparisons across multiple CNN architectures remain scarce. Furthermore, CNN architectures widely used for medical image classification carry high parameter counts that limit feasibility in resource-constrained settings, motivating architectures that achieve competitive accuracy with substantially fewer parameters. Here we propose a parameter-efficient deep learning framework for pneumonia detection based on transfer learning, evaluated across three CNN architectures representing distinct architectural families: EfficientNet-B0 with fine-tuning (proposed method), ResNet50, and DenseNet121, trained under identical conditions on the Kaggle chest X-ray dataset (5,863 images). Our method achieved 90% classification accuracy, outperforming both baselines while requiring 4.8x fewer parameters than ResNet50. To evaluate explainability, Grad-CAM, Grad-CAM++, and Score-CAM were applied across all three architectures and compared quantitatively using Intersection over Union against manually annotated lung segmentation masks, Insertion score, and Deletion score, with pairwise statistical validation via Wilcoxon signed-rank tests and Bonferroni correction. Findings show that classification accuracy and XAI explanation quality must be evaluated independently, and that the proposed parameter-efficient architecture offers a favorable trade-off for resource-constrained clinical deployment.
Jedamzik, T. A.; Martens, J.; Siebes, M.; van den Wijngaard, J. P. H. M.; Schreiber, L. M.
Show abstract
BackgroundQuantitative dynamic contrast-enhanced myocardial perfusion cardiovascular magnetic resonance (CMR) enables estimation of myocardial blood flow (MBF) and myocardial perfusion reserve (MPR). These measurements require an arterial input function (AIF), which is typically derived from the left ventricular blood pool. However, the contrast agent bolus undergoes dispersion during transport through the coronary vasculature before reaching the myocardial microcirculation. This may introduce systematic and spatially heterogeneous errors in MBF and MPR estimates. PurposeThis work provides an extended segmental analysis of bolus-dispersion-induced errors in quantitative myocardial perfusion MRI using previously established computational fluid dynamics (CFD) simulations in realistic porcine coronary artery models. The focus of the present analysis is the assignment of coronary outlets to myocardial segments and the resulting segmental variability of MBF and MPR errors. MethodsRealistic three-dimensional models of the left and right coronary arteries were extracted from an ex-vivo porcine imaging cryomicrotome dataset. The models extended down to the pre-arteriolar level and included 364 outlets for the left coronary artery and 104 outlets for the right coronary artery, with an average outlet diameter of 383 {+/-} 85 {micro}m. Blood flow was simulated under rest and stress conditions using OpenFOAM. Contrast agent transport was then modeled by solving the advection-diffusion equation using a gamma-variate bolus as input. Outlet concentration-time curves were analyzed using an indicator-dilution model to estimate MBF and MPR errors. Outlets were assigned to standardized myocardial segments, and segmental averages were evaluated with respect to coronary supply territory and travel distance from the model inlet. ResultsThe simulations demonstrated marked segmental heterogeneity of volume blood flow and bolus-dispersion-induced MBF and MPR errors. Errors increased with travel distance from the coronary artery inlet and were more pronounced in regions supplied by the right coronary artery, consistent with lower flow velocities and stronger bolus dispersion. The resulting systematic errors led to underestimation of MBF and overestimation of MPR, with segmental deviations reaching up to approximately 60%. ConclusionBolus dispersion in the coronary vasculature may lead to substantial segmental and location-dependent errors in quantitative myocardial perfusion MRI. This extended analysis indicates that dispersion-related bias is not spatially uniform, but depends on coronary supply territory, travel distance, and flow conditions. These effects should be considered when interpreting regional MBF and MPR estimates, particularly as automated quantitative myocardial perfusion CMR becomes more widely used.
Tzanis, E.; Klontzas, M. E.
Show abstract
This study presents ReCo (Research Cosmos), a self-configuring and self-extending agentic research framework for the biomedical domain. ReCo is orchestrated by a large language model that interacts with native computing tools, bundled Model Context Protocol (MCP) servers, structured skills, persistent project memory, and a desktop interface. Its bundled MCP servers provide biomedical analysis capabilities while serving as implementation paradigms for integrating new computational and AI frameworks. Structured skills encode procedures for environment configuration and framework ingestion, enabling ReCo to inspect repositories, manuscripts, or local codebases; identify dependencies and execution patterns; create isolated runtime environments; design and implement MCP interfaces. Self-extension was evaluated using five heterogeneous systems: the Merlin computed tomography foundation model, MAISI-v2 medical image synthesis framework, asari liquid chromatography-mass spectrometry workflow, DosimeTron agentic radiation-dosimetry platform, and Orthanc DICOM server. ReCo successfully operationalized all five systems and completed predefined functional evaluations. Re-hosted DosimeTron outputs demonstrated near-perfect agreement with the reference pipeline across 651 organ observations (Pearson correlation and Lin concordance correlation coefficient, 0.99999; mean absolute percentage difference, 0.37%). Notably, ReCo configured Orthanc as a PACS-like coordination layer, integrated it with DosimeTron, Merlin, and TotalSegmentator, and orchestrated data retrieval, analysis, and return of valid DICOM RTSTRUCT, RTDOSE, and Structured Report. ReCo provides a unified environment for configuring, documenting, and operationalizing heterogeneous biomedical frameworks, reducing technical barriers to the adoption and integration of emerging computational and AI methods. The official open-source ReCo GitHub repository is available at: https://github.com/eltzanis/ReCo
Sonoda, Y.; Yamagishi, Y.; Hirano, Y.; Miki, S.; Nakao, T.; Hanaoka, S.; Nomura, Y.; Hamada, A.; Kanemaru, N.; Miyo, R.; Takahashi, M. M.; Hosoi, R.; Yoshikawa, T.; Abe, O.
Show abstract
Purpose: To evaluate the latest open-weight vision-language models (VLMs) on the Japanese Diagnostic Radiology Board Examination (JDRBE), assessing overall accuracy and the effects of image input, reasoning, and language. Materials and Methods: In this retrospective study, 29 open-weight VLMs from 13 developers, released in or after January 2025, were evaluated on 327 image-bearing questions from four years of the JDRBE, a non-public benchmark with low risk of data leakage. Each question was answered by each model with and without the image(s), under three language conditions and with reasoning enabled and disabled. Accuracy was the primary outcome, and within-model differences were tested with paired bootstrap confidence intervals and sign-flip permutation tests with Benjamini-Hochberg correction. Results: In the Japanese condition with image input and reasoning, the leading models reached 73.7% (gemma-4-31B-it), 73.1% (Qwen3.5-397B-A17B), and 72.1% (Kimi-K2.6). On the 2025 subset, these three models (74.1%-75.5%) scored above the mean accuracy of five newly board-certified radiologists who passed the 2025 examination (72%; range, 65%-83%). Accuracy broadly scaled with model size, although compact gemma-4-31B-it matched larger models. Enabling reasoning improved accuracy in nearly all models and the contribution of image input was larger when reasoning was enabled, particularly in higher-performing models. English prompts generally outperformed Japanese prompts. Conclusion: Several open-weight VLMs, without medical adaptation, performed at or above the mean of newly board-certified radiologists on the JDRBE, with both model size and reasoning contributing. The highest Japanese-language accuracy came from a compact model suitable for parameter-efficient fine-tuning and serving on a single graphics processing unit.
Wang, J.; Tang, W.; Ma, X.; Yan, H. m.; Yuan, Y.
Show abstract
Large language models (LLMs) are increasingly used for automated quality control (QC) of radiology reports. However, the reliability of LLMs on reports in Mandarin, and the relative performance of domestic versus international flagship models, remain unknown. We benchmarked 14 LLM configurations, seven Chinese-developed ("domestic") and seven international models, on 1,000 whole-body 18F-FDG PET/CT reports split into an error-injected "junior-docto" arm and a low-residual "finalised" arm (500 each), using a controlled error-injection gold standard. Under each blinded zero-shot prompt, each model flagged six error types and assigned a 1-5 overall score. Two distinct abilities: error-detection macro-F1 (0.356-0.667) and overall-score calibration (ICC[2,1] 0.099-0.627), were weakly and not significantly correlated across models (Spearman {rho} = 0.38, p = 0.18); the dissociation was instead evident in sharp rank reversals, the strongest detector (Claude-Opus-4.8 0.667) calibrating poorly (0.491), while the three best-calibrated models were all domestic (MiMo 0.627, GLM-5 0.612, DeepSeek 0.609). Once the access channel was controlled, domestic and international error detection were statistically indistinguishable ({Delta}macro-F1= -0.011, P = 0.84); domestic models showed consistent but not significant advantages in calibration ({Delta}ICC = +0.142) and Chinese-character-error detection ({Delta}F1 = +0.109), accompanied with large reductions in cost (US$0.09-2.71 vs $0.26-14.5 per 1,000 reports) and on-premise deployability. Re-running two flagships through both agent channels and clean APIs showed that agent channel inflated both detection and calibration (GPT-5.5 {Delta}ICC = +0.098, 95% CI 0.070-0.128), confirming that uncontrolled benchmarks over-credit agent-channel models. Missed-diagnosis detection was the universal weakness (best 0.467) and the one category where the human physicians outperformed every model. Raw detection ability does not guarantee a trustworthy score, and domestic and international models differ by deployment-relevant profile rather than by overall performance rank; both essential distinctions for performing clinical nuclear-medicine QC.
Loyd, Y. M.; Chase, S. E.; Krendel, M.
Show abstract
Nephrons are the functional units of the kidney; within each nephron, the glomerulus is the initial site of selective filtration that allows removal of waste products while preserving proteins in the bloodstream. Each glomerulus consists of a network of capillaries surrounded by specialized epithelial cells, podocytes, which mediate selective filtration. Abnormalities in glomerular structure impair renal function, resulting in proteinuria and kidney disease. Although several microscopy-based approaches exist to characterize glomerular architecture and structural abnormalities, quantitative analysis is often limited by labor-intensive image segmentation. In this study we present a semi-automated approach for segmentation and analysis of glomerular architecture from three-dimensional confocal microscopy data. Using mTmG transgenic mice that express membrane-associated EGFP in podocytes and membrane-associated tdTomato across all other cell types, we reconstruct podocyte processes and glomerular capillaries from volumetric renal images. This semi-automated approach reduces manual segmentation effort and supports more efficient, standardized analysis of glomerular architecture in three-dimensional confocal microscopy datasets.
Naidu, J. S.; Baskaradoss, V.
Show abstract
Background: Artificial intelligence (AI), including generative and foundation-based methods, has rapidly expanded within medical imaging research. However, the structure, citation impact, collaboration patterns, and thematic orientation of national research ecosystems remain incompletely characterised. Objectives: To evaluate global research trends in AI applied to medical imaging between 2017 and 2025, with detailed analysis of United Kingdom (UK)-affiliated output, citation performance, collaboration structure, funding landscape, and thematic evolution, with emphasis on generative and foundation-based methodologies. Materials and Methods: A bibliometric analysis of Scopus-indexed publications (2017-2025) was performed using a predefined search strategy targeting AI and medical imaging concepts, with emphasis on generative and foundation-based terms. Records were analysed globally and filtered for UK affiliation. Descriptive indicators including total publications (TP), total citations (TC), citations per paper (CPP), and year-on-year growth were calculated. Co-authorship and keyword co-occurrence networks were generated using VOSviewer (v1.6.19). Results: A total of 13,452 publications were identified globally (194,650 citations; global CPP 14.47), of which 889 (6.61%) were UK-affiliated. The UK ranked fourth by publication volume yet demonstrated higher citation efficiency (CPP 21.00) than several higher-volume countries. UK output increased approximately 18-fold between 2017 and 2025, with evidence of a citation-lag effect in recent years. Research activity was concentrated within a small number of institutions accounting for nearly half of national output, although citation impact varied independently of volume. Journal-dominant dissemination was associated with higher average citation impact compared with conference-centric models. Keyword analysis identified three principal thematic clusters: generative/deep learning methodologies, MRI- and diffusion-focused applications, and broader diagnostic imaging workflows. Highly cited publications were initially dominated by generative adversarial network-based reconstruction and synthesis, with recent rapid citation growth observed in diffusion and foundation-model architectures. Conclusion: UK-affiliated research represents a rapidly expanding and highly cited component of the global AI medical imaging literature, with increasing emphasis on generative, diffusion-based, and foundation-model approaches. These findings provide a reproducible bibliometric baseline for monitoring research activity, collaboration patterns, and potential translational priorities, while recognising that citation-based indicators do not directly measure clinical implementation, methodological quality, or real-world impact.
Vu, J.; Khodabocus, I.; Derzi, S.; Henry, M.; Davidge, S. T.; Macala, K.; Bourque, S. L.; Noble, R. M. N.
Show abstract
Background Perioperative incidents such as hypoxic cardiac injury often have subtle or nonspecific clinical manifestations. Reduction in myocardial oxygenation precedes biochemical changes, as well as electrical and functional changes. Photoacoustic imaging (PAI) is a modality that uses laser irradiation of tissue to generate ultrasonic waves, enabling spatially resolved quantitative mapping of oxygenated and deoxygenated haemoglobin. We investigated the utility of PAI for real-time monitoring of myocardial and great vessel oxygenation. Methods Male CD-1 mice were anaesthetised, and photoacoustic and simultaneous B-mode images were acquired of the myocardium and right ventricular outflow tract (RVOT), the pulmonary artery, and aorta. PAI was performed at fractional inspired oxygen levels (FiO2) of 100%, 21%, and then 10%. Separate cohorts of mice were exposed to increasing intravenous doses of either combined phenylephrine and isoprenaline, or individual administration of vasoactive or adrenergic agents. Results PAI reliably distinguished changes in oxygenation in the RVOT cavity, pulmonary artery, aorta, and myocardium. PAI detected hypoxia-induced changes in oxygenation, revealing greater desaturation in the myocardium than in the RVOT (-9.85%, 95% CI -14.94 to -4.77, P<0.0001). Escalating doses of phenylephrine and isoprenaline caused a progressive desaturation of the myocardium and RVOT (mean [95% CI]; myocardium 16 mg/kg: -14.64% [-27.62 to -1.65], P=0.0038 and RVOT 32 mg/kg: -18.71% [-32.15 to -5.27], P=0.0003). Myocardial deoxygenation was detected before changes in systolic function or electrical abnormalities. Conclusions This work demonstrates that PAI can reliably monitor cardiac oxygen desaturation, potentially offering an earlier warning of cardiac dysfunction and injury compared to existing monitoring tools. Keywords: Echocardiography, hypoxaemia, hypoxia, myocardial injury, oxygenation, perioperative monitoring, photoacoustic imaging